The Lancet Digital Health
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match The Lancet Digital Health's content profile, based on 25 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Tamm, A.; Shine, B.; James, T.; Withers, J.; Salih, H.; East, J. E.; Oke, J.; Davies, J.; Morris, E. J.; Nicholson, B. D.
Show abstract
Background The faecal immunochemical test (FIT) is central to triaging symptomatic patients with suspected colorectal cancer (CRC) in UK primary care, yet only about one in eleven patients above the NICE 10 ug/g threshold have CRC. Existing prediction models attempting to improve on FIT have relied on conventional statistics and limited predictors. Methods GP-requested FITs with linked data (Jan 2017 - May 2025) were extracted from the Oxford University Hospitals (OUH) datawarehouse. Patients aged [≥]18 with core bloods and 180-day CRC follow-up were included. Machine learning (ML) models were trained on up to 1,025 predictors: FIT, age, sex, blood tests and their time series slopes, diagnoses/procedures/prescriptions, deprivation, BMI, and ethnicity. Models comprised penalised logistic regression, generalised additive models (EBM, NAM, SNAM, NODE-GAM), decision tree ensembles (random forests, XGBoost), and a multilayer perceptron. Referral reduction versus FIT [≥]10 ug/g was evaluated at model risk score thresholds capturing the same cancers (conservative) or same proportion of cancers (less conservative) as FIT. Potential to prioritise referred patients was assessed by examining whether positive predictive value (PPV) is very high (>30%) at any substantial sensitivity (>10%). Nested twice-repeated five-fold cross-validation provided unbiased estimates. An existing COLOFIT model was evaluated alongside. Findings 62,219 individuals (746 CRC) were analysed; 30,862 patients (315 CRC) with high/low risk symptoms and buffered FITs formed the primary subset. At [≥]10 ug/g, FIT had 91.4% sensitivity, 84.2% specificity, 5.6% PPV, and 99.9% NPV. No model reduced referrals when required to capture the same cancers as in the FIT [≥]10 ug/g cohort. Generalised additive models achieved up to 18.5% referral reduction when detecting the same proportion but some different cancers as FIT [≥]10 ug/g (EBM: 18.5%, NODE-GAM: 17.5%, SNAM: 17.4%, COLOFIT: 16.7%). At 30% sensitivity, EBM, NAM and NODE-GAM had average PPVs between 34.6%-35.0%, while FIT had a PPV of 14.6%. Interpretation Generalised additive models (GAMs) reduced referrals on average by 19% if a small proportion of the FIT-positive CRCs were substituted with originally FIT-negative CRCs by the models. No model, including COLOFIT, reduced referrals while capturing all FIT-positive cancers. Generalised additive models could detect about a third of CRCs faster, as one in three patients flagged by the models had CRC at 30% sensitivity. Funding EPSRC Centre for Doctoral Training in Health Data Science; National Institute for Health Research (NIHR) Oxford Biomedical Research Centre; Cancer Research UK. Keywords Colorectal cancer, faecal immunochemical test, machine learning, positive predictive value
Wang, J.; Tang, W.; Ma, X.; Yan, H. m.; Yuan, Y.
Show abstract
Large language models (LLMs) are increasingly used for automated quality control (QC) of radiology reports. However, the reliability of LLMs on reports in Mandarin, and the relative performance of domestic versus international flagship models, remain unknown. We benchmarked 14 LLM configurations, seven Chinese-developed ("domestic") and seven international models, on 1,000 whole-body 18F-FDG PET/CT reports split into an error-injected "junior-docto" arm and a low-residual "finalised" arm (500 each), using a controlled error-injection gold standard. Under each blinded zero-shot prompt, each model flagged six error types and assigned a 1-5 overall score. Two distinct abilities: error-detection macro-F1 (0.356-0.667) and overall-score calibration (ICC[2,1] 0.099-0.627), were weakly and not significantly correlated across models (Spearman {rho} = 0.38, p = 0.18); the dissociation was instead evident in sharp rank reversals, the strongest detector (Claude-Opus-4.8 0.667) calibrating poorly (0.491), while the three best-calibrated models were all domestic (MiMo 0.627, GLM-5 0.612, DeepSeek 0.609). Once the access channel was controlled, domestic and international error detection were statistically indistinguishable ({Delta}macro-F1= -0.011, P = 0.84); domestic models showed consistent but not significant advantages in calibration ({Delta}ICC = +0.142) and Chinese-character-error detection ({Delta}F1 = +0.109), accompanied with large reductions in cost (US$0.09-2.71 vs $0.26-14.5 per 1,000 reports) and on-premise deployability. Re-running two flagships through both agent channels and clean APIs showed that agent channel inflated both detection and calibration (GPT-5.5 {Delta}ICC = +0.098, 95% CI 0.070-0.128), confirming that uncontrolled benchmarks over-credit agent-channel models. Missed-diagnosis detection was the universal weakness (best 0.467) and the one category where the human physicians outperformed every model. Raw detection ability does not guarantee a trustworthy score, and domestic and international models differ by deployment-relevant profile rather than by overall performance rank; both essential distinctions for performing clinical nuclear-medicine QC.
Amiruddin, N.; Mellor, S.; Crisp, R.; Nair, A.; Patel, M.
Show abstract
Background Ventilator-associated pneumonia (VAP) is the most frequent nosocomial infection in critical care, affecting 20-36% of mechanically ventilated patients. Early prediction is hampered by the absence of a reliable, objective diagnostic standard. We developed ADVISE (Automated Dudley Ventilation Infection Series Evaluation), a machine learning model to predict physiological deterioration consistent with developing VAP using routinely collected electronic health record data from a UK NHS intensive care unit. Methods Retrospective observational study of admissions at Russell's Hall Hospital ICU (2008-2026). Following National Data Opt-Out exclusion (158 admissions, 4.2%), 3,566 admissions generated 33,208 candidate 48-hour observation blocks. Six temporal variables - FiO2, ventilator mode, P:F ratio, procalcitonin (PCT), secretion amount, and secretion description - were extracted across the baseline window (hours 1-24). A composite VAP-surrogate outcome required concurrent P:F ratio decline (>=5%) and PCT rise (>=0.5 ng/mL) across the outcome window (hours 25-48). After sequential quality filters, 2,134 blocks (18 positive, 0.84% prevalence) were retained. An XGBoost classifier was trained using nested 5-fold cross-validation with scale_pos_weight=114.0 and ROC-based hyperparameter optimisation on 1,495 training blocks, evaluated on 639 held-out test blocks. Performance was assessed via AUROC, AUPRC, and calibration (Brier score). Bootstrap resampling (1,000 iterations) generated 95% confidence intervals. Results On the held-out test set (n=639, 5 positive outcomes), ADVISE achieved AUROC 0.874 [95% CI: 0.771-0.939] and AUPRC 0.031 [0.008-0.069], representing a 4.0-fold improvement over the no-skill baseline. Nested cross-validation mean AUROC was 0.844 +/- 0.078 (range 0.716-0.915). At the Youden-optimal threshold, sensitivity was 0% with specificity 97.8%, reflecting extreme class imbalance (0.78% test prevalence). A threshold targeting 80% sensitivity achieved sensitivity 80.0% [33.3-100.0%], specificity 87.4% [84.8-89.9%], positive predictive value 4.8% [1.1-9.9%], and negative predictive value 99.8% [99.4-100.0%], detecting 4 of 5 VAP cases with approximately 80 false alarms (12.6% false positive rate). Brier score was 0.0078. Feature importance identified baseline P:F ratio as the dominant predictor (41.3% total gain), followed by ventilator mode (26.1%), secretion amount (13.2%), secretion description (9.1%), procalcitonin (5.9%), and FiO2; (4.5%). Conclusions ADVISE demonstrates that baseline oxygenation trajectory and ventilatory support patterns - derived exclusively from routinely charted ICCA variables - can identify admissions at risk of VAP-related physiological deterioration with meaningful discrimination (AUROC 0.874) despite severe class imbalance. The 80% sensitivity operating point offers a clinically actionable alert rate (12.6% FPR), supporting integration into existing ICU workflows. This proof-of-concept study establishes feasibility; multi-site prospective validation is required before clinical deployment.
Samuels, T. H.; Forrest-Hammond, R.; Stockford, C.; Harris, S. K.; Eyre, D. W.; Gupta, R. K.; Noursadeghi, M.
Show abstract
Background: Bacteraemia is associated with poor outcomes but the diagnostic gold standard, peripheral blood culture, takes up to 24 hours to become clinically actionable, hampering early management decisions in suspected infection. Single predictors and existing sepsis risk scores discriminate poorly, and few multivariable bacteraemia models have been adequately validated in UK populations. Methods: We developed a logistic regression model, using backwards AIC based selection of predefined candidate predictors routinely available within hours of hospital attendance, in a retrospective cohort of 33,874 hospital encounters at University College London Hospitals (UCLH) between 2019 and 2024. Continuous predictors were modelled using restricted cubic splines and missing data handled using multiple imputation. Model performance was assessed via internal external cross validation and prediction instability analysis, before temporal validation in held-out 2024 UCLH data and external validation in 53,669 hospital encounters from the Infections in Oxfordshire Research Database (IORD). Results: Bacteraemia occurred in 5.2% of UCLH and 8.9% of IORD encounters, respectively. Twenty predictors were retained, spanning demographics, comorbidities, vital signs and blood tests. Discrimination was stable across development time periods (pooled c-statistic 0.82, 95%CI 0.81 to 0.84) and was maintained in temporal (0.83, 0.79 to 0.87) and external validation (0.83, 0.82 to 0.83), with excellent calibration in external validation (calibration slope 1.08 (1.05 to 1.11); calibration-in-the-large 0.01 (-0.02 to 0.04)). The model outperformed single predictors, established risk scores, and a reconstructed comparator model, and showed superior net benefit in decision curve analysis. Performance was consistent across age, sex, ethnicity and socioeconomic subgroups but degraded when blood cultures were sampled more than six hours after attendance and varied by likely infection site. Conclusions: This model accurately predicts bacteraemia using routinely collected data available within hours of hospital attendance, with performance maintained in a large, independent external validation cohort. It offers a generalisable, clinically interpretable tool to support early decision-making in suspected infection, pending further work to establish optimal implementation thresholds.
Doeleman, T.; Brussee, S.; Valkema, P.; Kempf, W.; Vermeer, M.; Kers, J.; Wynaendts, L.; Kerckhoffs, K.; de Jonge, M.; Nguyen, A.; Peters, E.; Wobser, M.; Rauert-Wunderlich, H.; Rosenwald, A.; Stadler, R.; Jansen, P.; Battistella, M.; Roccuzzo, G.; Quaglino, P.; Schrader, A.
Show abstract
Background Histological diagnosis of early-stage mycosis fungoides (MF) is hindered by profound overlap with benign inflammatory dermatoses (BIDs), leading to diagnostic delays and extensive ancillary testing. We developed MIMIC (Multiple Instance-learning for Identification of Mycosis fungoides In Cutaneous biopsies), a weakly supervised deep learning model designed as a triage tool at initial H&E whole slide image (WSI) review to distinguish classic patch and plaque stage MF from BIDs. We externally validated the model and evaluated its clinical utility. Methods In this retrospective multicentre study, we trained a base model using weakly supervised attention based multiple instance learning on 3,339 WSIs from two Dutch centres. Crucially, all MF training labels were derived from a deeply phenotyped national cohort featuring strict multidisciplinary expert panel consensus diagnoses (the clinical gold standard). Transportability was evaluated on 371 WSIs from four independent European centres. A blinded reader study on 171 WSIs compared morphology only performance of MIMIC with 11 (dermato-)pathologists. We then retrained an updated model on all retrospective multicentre data and assessed clinical utility in a strictly held out, consecutive Utrecht cohort (2022-2023; 486 accessions, 863 WSIs). Primary analysis focused on classic MF versus BIDs (453 accessions). Decision curve analysis, using Platt scaled probabilities to correct for spectrum bias, evaluated net benefit at a prespecified, safety oriented threshold of 0.04. Findings The base model showed good multicentre transportability (mean centre specific AUROC 0.91; pooled AUROC 0.84). In the reader study, MIMIC achieved an AUROC of 0.87, exceeding the mean pathologist AUROC (0.79) and the best individual reader (0.83). In the consecutive MF versus BID cohort, the updated model achieved an AUROC of 0.87 (95% CI 0.81-0.92). At the 0.04 threshold, sensitivity was 97.8% (44/45 MF cases) and specificity 50.2%, reducing unnecessary ancillary workups by 39.9 per 100 screening cases versus a test all strategy. Interpretation By identifying nearly half of BIDs as low risk while preserving near complete sensitivity for classic early stage MF in a European digital pathology workflow, this unimodal H&E approach offers a scalable digital solution to reduce defensive ancillary testing and accelerate the diagnostic journey for patients with MF. Further validation is needed in non European centres and in populations with darker skin phototypes.
Gorenshtein, A.; Adiniaev, Y.; Omar, M.; Barash, Y.; Klang, E.; Daniel, O.
Show abstract
Background: The Glasgow Coma Scale (GCS) is a universal neurologic severity score in the intensive care unit and is incorporated into APACHE, SOFA, mortality prediction models, ICU benchmarking, and quality metrics. In mechanically ventilated patients, however, the verbal component cannot be assessed. Common conventions, including assigning a normal total GCS of 15 or excluding patients with missing verbal scores, may misclassify the sickest patients as neurologically normal or remove them from analysis. Objective: To quantify non-assessable verbal GCS examinations after acute brain injury and determine how different handling conventions affect severity scoring and mortality-model performance across two independent critical care databases. Materials and Methods: We conducted a retrospective cohort study of adults with acute brain injury during their first ICU stay in MIMIC-IV, with replication in eICU-CRD. A verbal examination was considered non-assessable when documented as No Response-ETT. We measured the burden and determinants of non-assessability, compared the MIMIC-IV derived GCS convention with a component-aware GCS, and evaluated mortality-model handling strategies. Results: Among 14,230 patients, 45.2% had a non-assessable verbal examination, and 47.5% of ventilated patients had no assessable verbal score in the first 24 hours. Non-assessability was strongly associated with mechanical ventilation and mortality. The MIMIC-IV derived GCS assigned a score of 15 to 42.9% of patients and placed 11.6% in the lowest severity category despite eye and motor findings consistent with GCS [≤]9. Complete-case handling excluded 28.5% of patients, who accounted for 50.2% of deaths. Similar distortions were observed in eICU-CRD/APACHE across 171 hospitals. Discussion: Default-to-normal scoring can make severely ill intubated patients appear neurologically normal, while complete-case analysis removes the highest-risk patients. Conclusion: Non-assessable verbal GCS in mechanically ventilated patients should be explicitly flagged and reported in ICU severity scores, risk-adjusted mortality models, and benchmarking systems.
Forrest-Hammond, R. W.; Gupta, R.; McVean, G.; Noursadeghi, M.; O'Grady, J.; Samuels, T. H.; Eyre, D. W.
Show abstract
Background Bloodstream infections are a major cause of mortality, yet the primary testing method, blood cultures, have low positivity (<10%) and turnaround times of 24 - 48 hours. Many are taken from patients at low risk of infection, while some bloodstream infections are diagnosed late or missed entirely. We aimed to develop and externally validate machine learning models to improve targeting of blood culture testing. Methods In this retrospective cohort study, we used routinely collected clinical and laboratory data available around culture collection from a large multi-site NHS trust (Oxford University Hospitals; Infections in Oxfordshire Research Database), between 1 January 2016 and 17 March 2025. All blood cultures taken from adults and children were included. XGBoost models were trained to predict pathogenic blood culture positivity using a temporal split (training before 1 January 2024; held-out test thereafter). External validation used emergency department data (between 1st May 2019 and 30th April 2024) from University College London Hospitals. An additional analysis examined blood culture reallocation towards the highest-risk untested admissions. Findings 294,064 cultures were included (positivity 5.6%). In the temporal hold-out test set (n=46,339), AUROC (Area Under the Receiver Operating Characteristic) was 0.853 (95% CI 0.846 - 0.860), rising to 0.876 in emergency department patients, and the model was well calibrated (slope 1.046). In external validation (n=37,326), AUROC was 0.847 (95% CI 0.839 - 0.856) with preserved calibration. In a simulated resource-neutral reallocation, replacing the 10,000 lowest-risk sent cultures with the highest-risk untested emergency admissions yielded 627 additional positive cultures (28.3% relative increase in yield). Performance was reduced when restricted to data available at the point of culture collection (AUROC 0.769, 95% CI 0.760 - 0.779). Interpretation An externally validated, well calibrated machine learning model built from broadly available, routinely collected data could improve blood culture yield without increasing testing volume, supporting resource-neutral diagnostic stewardship across NHS sites.
Guo, J.; Younis, Y.
Show abstract
Background: To develop and validate multiple Machine Learning (ML) algorithms that predict Mechanical Ventilation (MV) requirement in Guillain-Barre Syndrome (GBS), and to determine whether they outperform the additive, score-based prognostic models in current use. Methods: This retrospective study analysed 233 GBS patients (training set, n = 186; validation set, n = 47). Five algorithms (Deep Neural Network (DNN), Extreme Gradient Boosting (XGBoost), Logistic Regression (LR), Random Forest (RF), and Naive Bayes (NB)) were trained and compared. Predictors were chosen by a three-method consensus pipeline executed inside each nested cross-validation fold, retaining 11 features. Whether BorderlineSMOTE was applied was determined per model by Optuna hyperparameter tuning. Hyperparameter tuning, probability calibration, and bootstrap resampling were applied; performance used accuracy, recall, F1, specificity, AUROC, and Brier score, with SHapley Additive exPlanations (SHAP) for model interpretability. Results: XGBoost achieved the strongest clinical performance (AUROC 0.807, accuracy 0.787, and recall 0.857), exceeding the validated EGRIS for MV (AUROC = 0.62). Calibration preserved recall (0.857) and shifted the operating point by one false positive while lowering the Brier score from 0.210 to 0.110 (naive Brier baseline 0.127, BSS = 0.134), so the deployed tool was developed using the probabilities from the calibrated XGBoost model. Consensus selection retained eleven predictors; blood prealbumin, blood FT3, and NLR ranked highest by both embedded importance and SHAP. The model was deployed as an interactive prognostic tool predicting MV risk at admission. Conclusions: ML algorithms substantially improve GBS prognosis by integrating eleven biomarker predictors, modelling nonlinear relationships, and providing SHAP-based interpretability. The single-centre sample is small, so external validation in larger, multi-centre cohorts is required before clinical deployment.
Nam, Y.; An, T.; Hwang, S. I.; Jang, H.; Jeon, C.; Jeong, J. W.; Jeong, J.; Kim, D. Y.; Kim, S. Y.; Kim, S.; Kim, Y.; Lee, K. H.; Oh, H. S.; Park, J. H.; Seo, M.; Sim, Y.; Song, J. M.; Song, S.; Yoon, H. M.; the MeducAI Reader Study Group, ; Hong, P.; Kim, N.
Show abstract
Background: Frontier text-to-image models can synthesise radiologic images of high realism, raising the question of whether expert radiologists can serve as a provenance safeguard for the medical image record. Methods: We conducted a prospective, pre-registered visual Turing test in which 60 invited Korean board-certified radiology faculty and trainees judged authentic (teaching-repository) and AI-generated radiologic images from a locked pool of 241 displayable cells (82 entities; nine subspecialties; six modalities; 60 readers x 60 trials = 3,600 reader-image observations) produced by two contemporary commercial generators. The primary endpoint was the confidence-weighted, reader-averaged multi-reader multi-case area under the curve for AI versus authentic images, conditional on the locked image pool; the key secondary endpoint was the Faculty-minus-Junior difference under a two one-sided tests equivalence framework. The pre-specified statistical analysis plan was registered on the Open Science Framework before data lock. Findings: All 60 readers completed the test. The pooled confidence-weighted area under the curve was 0.71 (95% CI, 0.69 to 0.74), above the null value of 0.5 but within the pre-specified modest tier (0.60 to 0.75). The Faculty-minus-Junior contrast was 0.04 (95% CI, -0.02 to 0.10), including zero, and the two one-sided tests established equivalence within the +/-0.10 margin. No reader stratum and no pre-specified sensitivity analysis reached the deployable-classifier threshold (area under the curve >= 0.75). Interpretation: In this single-country cohort, expert radiologists distinguished frontier-generated from authentic radiologic images only modestly, without a meaningful expertise gradient (equivalence within +/-0.10) and with no reader stratum reaching a standalone provenance safeguard. These findings support radiology AI-literacy training and pipeline-level provenance safeguards rather than reliance on reader judgment, and warrant retesting in an independent reader cohort. Funding: This research was supported by a grant of the Korea Health Technology R&D Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (grant number: RS-2025-02213531).
Naidu, J.; Muralidharan, S.; Prashani, A.; Baskaradoss, V.
Show abstract
Objectives: To test whether radiology report evaluation metrics distinguish clinically meaningful errors from textual changes and align with radiologist-assessed error burden. Methods: Cross-dataset evaluation used ReXErr-v1 (2,708 report pairs; 5,724 paired error sentences) and 100 RadEvalX report pairs with expert error counts. BLEU-4, ROUGE-L and METEOR were assessed in ReXErr-v1; RadEvalX analyses included these plus BERTScore, CheXbert, RadGraph F1 and RadCliQ. Outcomes were ReXErr-v1 pairwise win rate and AUROC for clinical-content versus linguistic errors, and RadEvalX Spearman correlation with clinically significant error count and AUROC for any significant error. Confidence intervals used 10,000 clustered percentile bootstrap resamples; Holm adjustment-controlled multiplicity. Results: ReXErr-v1 paired-sentence win rates were 0.986 for BLEU-4, 0.999 for ROUGE-L and 0.998 for METEOR, but discrimination of clinical-content from linguistic errors was modest (AUROC 0.609-0.620). Penalty magnitude was strongly associated with textual change after adjustment for error type (normalised character edit distance coefficient 0.746; 95% CI 0.705-0.788; P<0.001). In RadEvalX, CheXbert showed the highest correlation with clinically significant errors (rho=0.413; 95% CI 0.223-0.578) and highest AUROC (0.742; 95% CI 0.638-0.836). Conclusions: Near-ceiling sensitivity to textual corruption did not imply sensitivity to clinical significance. CheXbert showed the highest alignment with expert error assessment, although pairwise superiority was not demonstrated over all comparators and performance remained moderate.
Gallego Luxan, B.; Huberts, L.; Yu, J.; Blake, V.; Liu, L.; Jorm, L.; Ooi, S.-Y.
Show abstract
Background: Unplanned emergency readmissions remain common following hospitalisation for heart failure (HF). Residual congestion, atrial fibrillation, frailty, and other comorbidities contribute to adverse outcomes after discharge. Identifying patients at high risk of readmission or death may help target post-discharge management. Methods: We conducted a retrospective cohort study of patients hospitalised with HF in selected New South Wales hospitals who were discharged alive and not documented as receiving end-of-life care. Clinical, laboratory, medication, and text-derived variables extracted from electronic health records were used to develop predictive models and corresponding risk scores for emergency readmission and all-cause mortality within 180 days of discharge. Feature importance methods were used to identify key predictors and explain individual risk estimates. To illustrate model predictions while preserving patient privacy, we generated representative synthetic patient profiles by summarising the characteristics of groups of patients with similar predicted risk patterns and visualised the major contributors to their predicted risks using Shapley values. Results: The study included 5,202 hospitalisations among 3,933 patients. Within 180 days of discharge, 45.2% of patients experienced at least one emergency readmission and 12.4% died. The most common causes of emergency readmission were recurrent HF, followed by atrial fibrillation, chest pain, and pneumonia. Predictive performance was moderate for emergency readmission (AUC 0.70; calibration slope 1.30) and good for mortality (AUC 0.84; calibration slope 1.01). Emergency readmission risk was primarily associated with greater prior healthcare utilisation, a higher number of active medical problems, high risk of falls, older age, and impaired kidney function. Mortality risk was most strongly associated with abnormal red blood cell distribution width, elevated blood urea, older age, and lower systolic blood pressure. A lower number of discharge medications, particularly cardiovascular therapies, was associated with a higher risk of emergency readmission and a lower risk of mortality. Representative synthetic patient profiles demonstrated heterogeneity in the factors contributing to predicted risks, illustrating the value of patient-level risk visualisation. Conclusions: Predictive models identified clinically meaningful predictors of emergency readmission and mortality following HF hospitalisation. Patient-level visualisation of individual risk drivers may support more personalised post-discharge management.
Hwang, Y.-M.; Cui, Y.; Xu, J.; Pan, T.; Li, R.; Rice, B.; Tian, L.; Hernandez-Boussard, T.
Show abstract
Comorbidity indices are widely used in clinical research to summarize disease burden. However, traditional indices were developed decades ago in limited populations using fixed weights that do not reflect the diversity of patients in modern healthcare. We present the Personalized Comorbidity Score (PCS), a data-driven framework for context-dependent comorbidity scoring designed to capture patient complexity while remaining accessible for broad research adoption. PCS was developed using Epic Cosmos, a large national EHR network encompassing over 8 million adult inpatient encounters from 2015 to 2020, with comorbidities defined using AHRQ Clinical Classifications Software Refined categories. Models were developed separately across eight age-sex subgroups using LASSO-penalized Cox regression for feature selection and restricted mean survival time for score derivation. PCS is available in two versions: PCS Core, incorporating age, sex, and comorbidities, and PCS Extended, which additionally incorporates socioeconomic and geographic variables. PCS Core and PCS Extended achieved AUROCs of 0.812 and 0.813 for one-year mortality, outperforming traditional indices (AUROC, 0.714-0.730). PCS demonstrated consistently lower subgroup calibration error across demographic and socioeconomic groups without including race or ethnicity as model features. PCS was further evaluated in two complementary external EHR datasets (Stanford Health Care and MIMIC-IV) with distinct patient populations and data structures, where it consistently outperformed traditional indices. Open-source R and Python packages are provided to support broad adoption. PCS provides an updatable framework for comorbidity measurement that is accurate, context-dependent, and designed to evolve alongside clinical practice.
Gorenshtein, A.; Adiniaev, Y.; Omar, M.; Barash, Y.; Klang, E.; Daniel, O.
Show abstract
Objective: To quantify the burden, structure, and downstream analytic consequences of "Unable to Assess" (UTA) delirium documentation in the intensive care unit (ICU). Design: Retrospective cross-sectional and repeated-measures study. Setting: A single US academic medical center (Medical Information Mart for Intensive Care IV [MIMIC-IV], 2008-2019). Patients: 72,944 adult ICU stays with at least 1 delirium screen. Interventions: None. Measurements and Main Results: Among 610,632 screens, 130,455 (21.4%; 95% CI, 21.0%-21.8%) were recorded as UTA, exceeding the 119,052 (19.5%) scored positive. The UTA fraction rose from 2.0% at a Richmond Agitation-Sedation Scale (RASS) score of 0 to 97.8% at RASS -4; 22.0% of UTA screens occurred in arousable patients, where UTA was associated with mechanical ventilation (odds ratio [OR], 3.43; 95% CI, 3.17-3.71) and non-English primary language (OR, 3.74; 95% CI, 3.43-4.08). Building the delirium label three ways from the same patients shifted prevalence modestly (32.1% to 30.8%) and prediction (area under the curve, 0.737 to 0.719) but most affected the delirium-mortality association: in a baseline-adjusted model the OR was 4.12 (95% CI, 3.88-4.36) under complete-case handling and fell to 2.16 (95% CI, 2.06-2.27) when UTA was recoded as negative. UTA was recoverable from the observed clinical state (area under the curve, 0.95). Conclusions: In this ICU cohort, Unable to Assess was the most common recorded delirium result other than Negative, exceeding positive screens; recoding it as negative roughly halved the apparent delirium-mortality association by relabeling deeply sedated, high-mortality patients. Delirium datasets should preserve and report UTA, whose concentration among arousable non-English-speaking patients is a measurable equity target.
Ebbert, J. L.; Perry, A.; Szymanski, J.; Della Corte, D.
Show abstract
Background: Deep-learning systems for Gleason grading are developed almost entirely on high-end clinical scanners and on cohorts from a small number of Western institutions, yet deployment increasingly involves other devices and other populations. These two distribution shifts, device and population, are rarely tested together on the same physical slides. The PAR dataset, from Erbil, Iraq, digitizes each biopsy on three scanners and provides three distinct pathologist grades, so it permits both tests at once on a Middle Eastern cohort. A concurrent study by the dataset originators validated a task-specific model and two foundation models on PAR; we complement it by testing an inde-pendently developed detect-then-grade pipeline and by separating scanner effects on detection from scanner effects on grading. Methods: We applied one fixed model de-veloped on North American and European material to all 1017 whole-slide images (339 slides from 185 patients, three scanners; 49.6% clinically significant cancer) with no scanner-specific or population-specific tuning. We measured cancer detection (area under the ROC curve of the predicted cancer-tissue fraction), all-slide ISUP agreement of the deployed detect-then-grade pipeline (quadratic-weighted kappa, QWK), and grading agreement on pathologist-confirmed cancers, at the slide level and, because a case carries up to two slides, at the patient level. The reference reader was S.A.; thresholds and operating points were cross-validated leave-one-out; scanners were compared by paired within-biopsy bootstrap and confidence intervals confirmed by patient-cluster bootstrap. Results: Detection was statistically equivalent across scanners (AUC 0.987 to 0.991; paired differences at most 0.003) and transferred to this non-Western cohort with no per-population tuning. At a 95% sensitivity operating point the deployed pipeline reached cross-validated all-slide QWK of 0.86, 0.81, and 0.86 (Grundium, Hamamatsu, Leica), matching the inter-pathologist ceiling of 0.81, against 0.23 to 0.62 for the ungated model. Grading of confirmed cancers was scanner dependent: the compact Grundium (0.63) did not differ from the clinical Leica (0.67; paired difference 0.04, 95% CI -0.03 to 0.11), while both exceeded Hamamatsu (0.44). Results held at the patient level, with grading somewhat lower for two scanners; the two slides of a case disagreed in grade in 43% of cases, and patient clustering did not widen the intervals. Conclusions: Can-cer-tissue fraction is a triage signal robust across scanner and transferable to an un-derrepresented population for detection, while grading is the scanner-sensitive step. Prostate grading models should be deployed as a detect-then-grade pipeline, with grading validated per device and confirmed on the local population.
Sanchez-Valle, J.; Zambrana, C.; Navarro-Martinez, A.; Costa, F. X.; Rocha, L. M.; Cirillo, D.; Violan, C.; Valencia, A.
Show abstract
Multimorbidity is the dominant clinical reality of primary care, yet the temporal dynamics governing when and how persistent comorbidity associations emerge remain poorly characterised. Most large-scale comorbidity studies adopt a single observation window after an index diagnosis, implicitly assuming that associations detectable at one year are equally detectable at five. Using 11 years of electronic health records from 5,821,197 individuals in Catalan primary care, we applied a matched cohort design across nine complementary follow-up windows, five cumulative (0-1 to 0-5 years) and four conditional (1-2 to 4-5 years), to 1,315 index diseases, identifying 144,030 significant directed comorbidity associations in the five-year network. We found that 60.1% of these associations required at least three years of follow-up and were undetectable in shorter-window analyses, demonstrating that observation window length is a primary determinant of which comorbidities can be observed. To organise this temporal heterogeneity, we introduce the biological clock of multimorbidity: a two-dimensional framework that positions ICD-10 disease categories according to their rates of cumulative signal attenuation and the persistence of conditional risk. This framework identifies four reproducible temporal patterns (episodic, chronic stable, chronic progressive, and transient-persistent) that are robust under bootstrap resampling, leave-one-disease-out sensitivity analysis, and alternative clustering approaches. The biological clock is systematically modulated by sex, with Blood/Immune and Musculoskeletal disorders showing the largest sex differences in temporal dynamics. Network analysis identified 19 disease "initiators" that generate broad downstream comorbidity burdens and 21 "sinks" representing convergent endpoints of multiple disease trajectories. Comparison with hospital-based Danish data from 6,909,676 individuals showed that shared associations were 2.7-fold enriched over chance expectation (hypergeometric test, p<10-300) and showed moderate concordance of effect sizes (Spearman {rho}=0.460), confirming that the comorbidity structure identified here reflects genuine, generalisable signal; nonetheless, only 3.6% of primary care associations were replicated in the hospital network, indicating that the two settings capture largely complementary segments of the disease co-occurrence landscape. Together, these findings establish the observation window length as a principal design parameter in EHR-based multimorbidity research and the biological clock as a framework for understanding how and over what timescale disease associations emerge, persist, and resolve.
Kudrot, N.; Si, Y.; Sanjaya, J.; Pathak, S.; Haghi, M.; Alaei, K.; Placencia, G.; Pishgar, M.
Show abstract
ICU trauma patients are clinically heterogeneous, and early mortality risk stratification may support monitoring and resource allocation. We developed machine learning models for 30-day mortality prediction using information recorded during the first 24 hours after ICU admission. In MIMIC-III, six feature configurations were trained using 3,411 patients and compared in a patient-level configuration-selection hold-out subset of 853 patients. The selected XGBoost configuration yielded an area under the precision-recall curve (AUPRC) of 0.556 and an area under the receiver operating characteristic curve (AUROC) of 0.863. For cross-database evaluation, a 228-predictor harmonized XGBoost model was refitted on the complete MIMIC-III cohort and evaluated in 13,747 MIMIC-IV ICU stays without using MIMIC-IV outcomes for model development or recalibration. It achieved an AUPRC of 0.495, an AUROC of 0.825, and a Brier score of 0.109. Calibration was monotonic but showed increasing overprediction at higher predicted risks. First-24-hour clinical information retained predictive value across MIMIC database versions, although internal configuration selection, model differences, same-center provenance, and incomplete feature-mapping documentation limit generalizability and deployment readiness.
Shrestha, L.; Maharjan, D.; Bista, U.
Show abstract
Objectives: To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal. Design: Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the STARD 2015 checklist and STARD-AI/DECIDE-AI guidelines. Setting: Department of Radiology and Imaging, Patan Academy of Health Sciences / Patan Hospital, Lalitpur, Nepal. Participants: 826 consecutive health assessment applicants (foreign employment Pre-Departure Medical Examination and student migration) undergoing chest radiography between 5 June and 20 June 2026. Two cases were excluded due to DICOM technical failure. Index test: DenseNet-121 algorithm (TorchXRayVision library, densenet121-res224-all pretrained weights). A maximum aggregated pathology probability score was derived per radiograph and compared against a post-hoc derived threshold of 0.6258 (selected as the highest threshold achieving the pre-specified >=95% sensitivity criterion). Reference standard: Single-reader-per-case review by one of three radiologists - two board-certified radiodiagnosticians (LS: 276 cases; DM: 275 cases) and one radiology resident (UB: 275 cases) - each blinded to AI output, using a standardised data collection worksheet capturing binary classification (abnormal/normal) and free-text findings. Results: Of 826 radiographs, 41 (4.97%) were classified as abnormal by the reference standard. At the post-hoc derived threshold of 0.6258, the DenseNet-121 algorithm achieved: sensitivity 95.12% (95% CI 83.9-98.7%), specificity 77.2% (95% CI 74.1-80.0%), area under the receiver operating characteristic curve (AUROC) 0.9583 (95% bootstrap CI 0.9225-0.9843), NPV 99.67% (95% Wilson CI 98.8-99.9%), PPV 17.89% (95% Wilson CI 13.4-23.5%), and Cohen's {kappa} 0.237 (95% bootstrap CI 0.174-0.304). Brier score was 0.3621 (null Brier 0.0472) and ECE was 0.564, confirming calibration failure due to score compression (range 0.52-0.72) despite preserved discrimination. Cross-validated results: Ten-fold cross-validation yielded bias-corrected sensitivity 95.12% (95% Wilson CI 83.9-98.7%; optimism 0.00 pp) and specificity 75.80% (95% Wilson CI 72.7-78.7%; optimism +1.40 pp), confirming primary metrics are not materially inflated by circular optimisation. Conclusions: The DenseNet-121 algorithm demonstrated high sensitivity and excellent discrimination for chest radiograph triage in a Nepali health-assessment population, supporting its potential as a rule-out tool (NPV 99.67%). Systematic score compression - preserved discrimination despite calibration shift - is a quantifiable marker of LMIC distributional shift. Prospective local calibration studies are warranted before operational deployment.
Kang, Y.-J.; Jun, S.-Y.; Kim, S.
Show abstract
Background. Breast cancer treatment depends on histopathological features, such as grade and receptor-defined subtype; however, specialist pathologist access is constrained when the workforce is limited. Commercial multimodal large language models (MLLMs) accept hematoxylin and eosin (H&E) image tiles through paid interfaces without local hardware or fine-tuning. However, prior pathology evaluations addressed only coarse tasks. Whether they reach treatment-determining accuracy and whether vendors agree remain unclear. Methods. We aimed to evaluate three vendor-designated flagship MLLMs (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.5) in 427 invasive breast cancer cases. Each case went to all three with identical H&E tiles and prompts, and the subtype was inferred in the second call. The reference was an institutional sign-out report of an immunohistochemistry-derived subtype. We calculated the concordance, sensitivity, specificity, Cohen's kappa, and pairwise McNemar and Bowker tests. Findings. Claude ranked highest by raw histologic-type concordance but lowest by kappa, classifying all 23 lobular and seven micropapillary carcinomas as invasive breast carcinoma of no special type. The models anchored the Nottingham grade to three modal grades. None of the models reliably identified human epidermal growth factor receptor 2-positive disease. The failure direction was vendor-specific: Claude and GPT-5.5 were under-detected, whereas Gemini was over-called. Twelve prompt variants (4,056 calls) did not recover sensitivity. Interpretation. No current commercial MLLM reaches deployment-ready accuracy for any treatment-determining feature of breast pathology. As each vendor fails in its own fixed direction, changing vendors alters the type of error rather than removing it; therefore, the value of these models is assistive rather than autonomous. At USD 0.20-0.50 per case, they may serve as supervised draft generators that leave the diagnosis with the pathologist.
Amagai, S.; Liao, W.-T.; Murphy, C.; Reamer, C.; Liu, Y.; Ambil, B.; Fernandes, G.; Santhosh, L.; Lyons, P.; Jordan, N.; Liebovitz, D.; Kline, A.; Rojas, J. C.; Luo, Y.; Gao, C.
Show abstract
ICU-to-ward transfers are high-risk transitions marked by information loss and burdensome handoff preparation. We developed PAUSE-Agents, a clinician-in-the-loop multi-agent LLM pipeline that drafts source-attributed handoff briefs from structured ICU data and clinical notes using the clinician-developed ICU-PAUSE template. Mirroring ICU team structure, PAUSE-Agents routes each record through a scribe extractor, 6 role-specialized agents, explicit conflict surfacing, and deterministic safety checks before synthesis, producing an editable first draft rather than an autonomous note. In a single-center medical ICU cohort, 5 physicians completed 100 reviews of 84 agent-drafted briefs. Among adjudicable claims, 98.8% were verified and 1.2% were incorrect; 88% of briefs had no pertinent omission, and mean PDSQI-9 quality was 4.20/5. PAUSE-Agents surfaced 118 conflict warnings and 421 safety flags, making documentation inconsistencies visible before handoff. An o4-mini PDSQI-9 judge showed limited case-level discrimination but supported aggregate monitoring. We release PAUSE-Agents and its clinician evaluation application.
Singh, P.; Platt, S.; Bussey, O.; Heacock, L.; Verdone, A.; Chen, W.; Reynolds, H. R.; Yu, C.; Shen, Y.; Bredella, M. A.
Show abstract
Purpose: To develop and evaluate a deep learning model for automated quantification of breast arterial calcification (BAC) on screening mammography and to assess whether AI-derived BAC burden predicts major adverse cardiovascular events (MACE) in women. Methods: In this retrospective study, 202,006 women who underwent screening mammography without history of MACE were included. A BAC segmentation model was trained on an expert-annotated dataset using a multi-task U-Net with a ResNet-18 encoder to detect and segment BAC. BAC burden was quantified as area (mm{superscript 2}) from model-generated masks using DICOM pixel spacing and categorized by tertiles into low, intermediate, and high. The PREVENT score and incident MACE were identified from electronic health records. Cox proportional hazards models were developed to evaluate AI-derived BAC burden and PREVENT score alone, and combined models for 5 - and 10-year cardiovascular risk prediction. Results: Among 202,006 women (mean age 54.8{+/-}11.7 years), 23.1% had AI-detected BAC, and 7,701 (3.8%) developed incident MACE during a median follow - up of 7.5 years. On the geographically held-out test set, the BAC model achieved an AUROC of 0.97, Dice score of 0.6678, and Pearson correlation of 0.961 between AI-derived and manually annotated BAC burden. BAC burden increased with age and was higher among women who developed MACE. Five - year MACE incidence increased across BAC categories from 1.5% in women without BAC to 6.9% in those with high BAC burden. BAC burden alone showed modest prediction of MACE, with 5-year and 10-year AUROCs of 0.661 and 0.650, respectively, while PREVENT achieved AUROCs of 0.781 and 0.771. Adding BAC to PREVENT produced minimal improvement in discrimination. Conclusion: Deep learning-based BAC quantification from routine mammography is feasible, accurate, and associated with future cardiovascular risk. Although BAC added little to PREVENT for overall discrimination, it may serve as a scalable opportunistic imaging biomarker to identify women at elevated cardiovascular risk and support preventive care.